Back

Journal of Dental Research

SAGE Publications

Preprints posted in the last 7 days, ranked by how well they match Journal of Dental Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Evaluating GPT-4o Model Proficiency and Clinical Reasoning for Antimicrobial Stewardship in Dentistry

Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.

2026-09-03 dentistry and oral medicine 10.64898/2026.09.01.26361980 medRxiv
Top 0.1%
4.6%
Show abstract

Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.

2
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 0.1%
2.1%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

3
EcoEnamel: Development of a Gelatin-Pectin Film for S. mutans Inhibition and Enamel Preservation in an In Vitro Model

Merle, J. A.; Javelona, G.

2026-09-01 microbiology 10.64898/2026.08.18.745620 medRxiv
Top 0.2%
0.6%
Show abstract

Rinsing-dependent dental hygiene presents a significant public health challenge in water-scarce environments. This study investigated combinations of xylitol (Xyl), chitosan (Chi), glycyrrhizin (Gly), epigallocatechin gallate (EGCG), dicalcium phosphate (DCP), and nano-hydroxyapatite (nHA) on the primary bacteria behind dental caries, S. mutans. These combinations were assessed for markers of dental caries by biofilm reduction, bacterial killing, and acid buffering against S. mutans when applied to an in vitro simulated enamel model using glass bead surfaces for biofilm formation, and gene expression was subsequently examined via RT-qPCR. Separately, mineral retention was also quantified. The EGCG-DCP-Xyl film demonstrated the highest overall efficacy, achieving a significant reduction in biofilm concentration compared to the untreated control and performing similarly in magnitude to the positive toothpaste control. Dead fluorescence staining confirmed that the EGCG-DCP-Xyl film induced the highest rate of non-viable cells, followed by the Chi-Gly film and the Gly-Xyl film. During 10-day pH cycling, the EGCG-DCP-Xyl and DCP-Xyl formulations buffered pH the most, consistently maintaining mean pH levels safely above the demineralization threshold of pH 5.5. The EGCG-DCP-Xyl also optimized mineral stability with the highest retained calcium concentration, significantly outperforming the Chi-Xyl film. At the transcript level, the EGCG-DCP-Xyl film induced substantial downregulation of key virulence genes, yielding decreases in expression for glucosyltransferase B (gtfB), associated with biofilm synthesis, collagen-binding protein (cnm), associated with tissue invasion, and lactate dehydrogenase (ldh), associated with lactic acid production, compared to the untreated control, with effects comparable in magnitude to the positive toothpaste control. This research suggests that targeting bacterial pathways and mineral loss through a portable film may have potential for preventing dental caries, especially in environments where water is limited. However, additional studies are necessary to evaluate real-world effectiveness.

4
Collagen staining with fast green FCF enables 3D imaging of pulmonary fibrosis

Saqib, M.; Rivers, A. K.; Masala, S.; Baker, J. R.; Hobbs, C.; Boden, A.; Jose, A. A.; Herzog, D.; Cleary, S. J.

2026-08-31 pathology 10.64898/2026.08.27.747478 medRxiv
Top 0.4%
0.2%
Show abstract

Current approaches for imaging fibrotic remodeling have sensitivity, specificity and cost drawbacks that limit both preclinical research and clinical diagnosis. Here, we show that fast green FCF, a small molecule that binds to fibrillar collagen, enables highly sensitive and specific imaging of fibrosis in lung samples from mice and humans using fluorescence microscopy. We report strategies for using fast green FCF staining to assess fibrotic remodeling using precision-cut lung slice and whole-biopsy preparations. Our findings demonstrate that fluorescence imaging of fast green FCF-stained collagen will be useful for fibrosis research and may help to improve detection of fibrosis in clinical pathology.

5
A Tunable Flat-Jet Hydro-Debridement Device: Clinical Feasibility for Soft Tissue Wound Management

DATTA, A.; Majumder, R.; Biswas, I.; Ganguly, R.; Santra, A. K.; Sarkar, S.; Gumta, M. K.; Sarkar, S.

2026-09-04 surgery 10.64898/2026.09.01.26361120 medRxiv
Top 0.8%
0.1%
Show abstract

Background: Chronic wounds, ulcers, and lacerations require staged debridement and irrigation to promote healing. However, conventional techniques of debridement, such as surgical, chemical, or autolytic, struggle to fully remove residual necrotic tissue, slough, and unhealthy granulation from wound sites, especially when lodged within wound clefts and cavities, and in wounds with exposed structures. This promotes polymicrobial biofilms, delays wound closure, and causes significant discomfort with increased morbidity. Objective: To demonstrate the feasibility of using an indigenously developed tunable flat-jet hydro-debridement device (presently termed as CleanseJet), a frugal wound debridement system designed for deployment in resource-constrained clinical settings. Methods: An open-label, interventional, single-centre, parallel-group pilot randomized controlled trial was conducted to clinically evaluate an indigenously developed tunable flat-jet hydro-debridement device in patients with wounds of varied aetiology. The device provided adjustable spray impact force and coverage area tailored to wound characteristics. Outcomes were compared with a control group receiving standard wound care alone, with time to complete granulation serving as the primary healing endpoint. Outcomes were compared with a control cohort receiving standard of care alone. Results: The removal of loose devitalized tissue, slough, and biofilms from the wound bed improved the healing, which were monitored using the SINBAD scoring system. No adverse events were reported, supporting the feasibility and safety of CleanseJet. Conclusion: While commercial hydro-debridement systems are effective, they are often costly, rely on disposable components, and require specialized training. In contrast, CleanseJet provides a low-cost, easy-to-use alternative that can be operated with minimal training, making it suitable for broader clinical use without observed adverse effects.

6
Sex Differences in the Impact of Allosensitization on Waitlist Access and Post-Transplant Outcomes in Adults with Congenital Heart Disease

Joseph, A.; Kearney, K.; Henricks, C.; Morgan, J. L.; Tan, W.; Shafer, K.; Wrobel, C.; Lacelle, C.; Burns, K.; Jawaid, A.; Tapaskar, N.; Solmonson, A.; Nelson, D. B.; Truby, L. K.

2026-09-02 transplantation 10.64898/2026.08.31.26361832 medRxiv
Top 1%
0.0%
Show abstract

Background: Adult congenital heart disease (ACHD) patients are prone to HLA-antibody formation from multiple surgeries, transfusions, and prosthetic surgical material. Females with ACHD may accrue additional, non-surgical alloantigen exposure. Whether sex modifies the impact of allosensitization on heart transplant (HT) access and outcomes in ACHD remains unknown. Methods: We retrospectively analyzed the OPTN/UNOS registry of adults with ACHD listed for first-time HT (2018-2025). Sensitization was defined by calculated panel reactive antibodies (cPRA) at listing. We tested the sex x sensitization (highly sensitized, cPRA >50%) interaction on transplant access using Fine-Gray competing-risks regression, treating transplantation as the event of interest and death or removal from the waitlist as competing events, and on post-transplant survival using multivariable Cox proportional-hazards regression, both adjusted for age at listing, mechanical support at listing, and the number of distinct prior cardiac surgery categories. Results: Among 856 candidates (38% female), females were more often highly sensitized than males (23% vs 14%; age-adjusted OR 1.81, 95% CI 1.26-2.61), even after adjusting for surgical burden. Sensitization reduced transplant access in females (84% to 71%; median wait 60 to 110 days, p < 0.001) but not males (79% vs 79%, median wait 88 vs 98 days). In adjusted Fine-Gray models, the subdistribution hazard for transplant was reduced in sensitized females (sHR 0.54, 95% CI 0.41-0.72) with no effect in males (sHR 0.96, 95% CI 0.73-1.26), and the sex x sensitization interaction was significant (interaction sHR 0.64, 95% CI 0.44-0.94, p = 0.02). Post-transplant mortality was numerically higher in sensitized than non-sensitized candidates in both sexes and the sex x sensitization interaction on 1-year mortality was not significant. The sex-asymmetric effect persisted and was more pronounced in the multiorgan candidates. Conclusions: Allosensitization is not a sex-neutral barrier to transplant in HT candidates with ACHD. Females are more sensitized and have reduced transplant access without differences in 1-year mortality. The female excess in sensitization is not accounted for by surgical burden, and the exposures responsible remain to be defined. These findings warrant a sex-aware listing strategy and further studies.

7
Burden of fatigue in compensated chronic liver disease: findings from the multinational a:GAP Study

Choudhuri, G.; Akhundova-Unadkat, G.; Naidoo, N.; Morales-Castillo, M.; Guillaume, X.; Duijnhoven, R. G.; Safaei, A.; Swain, M. G.

2026-09-02 gastroenterology 10.64898/2026.08.28.26361618 medRxiv
Top 1%
0.0%
Show abstract

Background & Aims: Fatigue is a central symptom of chronic liver disease (CLD), substantially impacting health-related quality of life (HRQoL). This study aimed to further understand CLD symptomatology, including fatigue, and its impact on HRQoL from a patient perspective. Methods: Abbott Global Assessment of Patients unmet needs (aGAP) was a multinational, cross-sectional survey in adults with compensated CLD in China, India and Mexico, conducted between July and November 2024. Adult participants who self-reported that they had physician-diagnosed CLD and were experiencing fatigue completed a quantitative survey to assess symptom burden and included three HRQoL patient-reported outcome (PRO) questionnaires (Patient-Reported Outcomes Measurement Information System [PROMIS]-29+2, Work Productivity and Activity Impairment - Specific Health Problem version 2.0 [WPAI: SHP], Multidimensional Fatigue Inventory [MFI]). Results: Overall, 505 participants (China: 200; Mexico: 105; India: 200) completed the study. Participants reported that their CLD-related fatigue sometimes, often or always affected their self-esteem/confidence (45.1%) and ability to maintain or acquire new employment (38.6%). Most participants reported moderate (51.3%) or serious (26.9%) fatigue, with 33.5% experiencing fatigue every day or almost every day. Many participants felt their social life was negatively impacted by their fatigue (47.3%) and that there were related financial difficulties (53.9%). Use of validated PRO tools demonstrated severe fatigue (MFI: overall mean [SD] 13.9 [3.4] general fatigue and 13.4 [3.6] physical fatigue) as well as substantial levels of work and activity impairment (WPAI: SHP overall mean [SD] 53.0 [26.4]) and high levels of anxiety, pain interference, depression and sleep interference (PROMIS T-scores [&ge;]54). Conclusions: Fatigue has a substantial impact on HRQoL among adults with CLD across several countries, highlighting a global unmet need for targeted interventions to effectively identify and manage the condition.

8
Evaluating Mean Platelet Volume in relation to Disease Severity in Paediatric Sickle Cell Anaemia: A Cross-Sectional Study in Kwara State, North-Central Nigeria

Oladimeji, F. D.; Adewoyin, A. D.; Oyeleke, K. O.

2026-09-02 hematology 10.64898/2026.08.28.26361349 medRxiv
Top 1%
0.0%
Show abstract

Background: Sickle cell anaemia (SCA) is characterised by chronic haemolysis, inflammation, platelet activation, and recurrent vaso-occlusive complications. Mean platelet volume (MPV) is a readily available platelet index, but evidence regarding its relationship with disease severity in paediatric SCA remains limited and inconsistent, particularly in African populations. Objective: To evaluate the relationship between MPV and disease severity among children with SCA in Kwara State, North-Central Nigeria. Methods: This hospital-based cross-sectional study included 51 clinically stable children with confirmed SCA consecutively recruited from the paediatric haematology clinic of Children Emergency Specialist Hospital, Ilorin. Complete blood count, including MPV, was performed using a Rayto RT-7600 automated haematology analyser. Disease severity was assessed using a composite clinical and laboratory scoring system based on a previously described method. Pearson's correlation, Spearman's rank correlation, simple linear regression, and the Kruskal-Wallis test were used as appropriate. Statistical significance was set at p < 0.05. Results: Of 51 participants, 14 (27.5%) had mild, 33 (64.7%) moderate, and 4 (7.8%) severe disease. Mean MPV was 9.34 +/- 0.76 fL (range, 8.0-11.2). Pearson's correlation showed a weak positive, non-significant linear relationship with severity score (r = 0.231, p = 0.103), whereas Spearman's analysis showed a weak positive monotonic association (rho = 0.286, p = 0.042). Regression explained 5.3% of severity-score variation (R2 = 0.053, p = 0.103). MPV did not differ significantly across severity categories (H = 2.163, p = 0.339). MPV correlated inversely with haemoglobin (r = -0.556, p < 0.001) and positively with platelet count (r = 0.307, p = 0.029). Conclusion: MPV showed a weak relationship with disease severity but inconsistent statistical evidence across analyses. The limited explained variance and absence of significant differences between severity categories do not support MPV as a standalone severity marker. Larger longitudinal studies are warranted. Keywords: Sickle cell anaemia; Mean platelet volume; Disease severity; Platelet indices; Paediatric haematology; Cross-sectional study; Nigeria.

9
Acute Renal, Hepatic, Thromboembolic and Functional Complications after Community-Acquired Acute Lower Respiratory Tract Infection: A Prospective Cohort Study in Bristol, UK, 2022-2024

Chatzilena, A.; Hyams, C.; Challen, R.; Lahuerta, M.; McGuinness, S.; Clout, M.; Begier, E.; King, J.; Morales-Aza, B.; Duale, K.; Rodriguez Pereira, A.; Healy, W.; Southern, J.; Wells, P.; Lihou, K.; Grimes, C.; Campling, J. A.; Maskell, N.; Oliver, J.; Vyse, A.; Gessner, B.; Finn, A.; Danon, L.; The AvonCAP Research Group,

2026-09-02 respiratory medicine 10.64898/2026.08.28.26361617 medRxiv
Top 1%
0.0%
Show abstract

Introduction Acute lower respiratory tract disease (aLRTD) is a leading cause of hospitalisation and death, particularly in older adults and adults with comorbidities, with acute lower respiratory tract infection (aLRTI; pneumonia and non-pneumonic LRTI) being a major component. Non-pulmonary complications and functional decline after aLRTI are recognised, but their pathogen-specific burden is poorly described. We aimed to quantify renal, hepatic, thromboembolic and functional complications, and mortality, after aLRTI hospitalisation, by clinical phenotype and pathogen. Methods We conducted a cohort study of adults (>18 years) admitted with aLRTD to two hospitals in Bristol, UK (01 August 2022-31 July 2024). aLRTD was classified as pneumonia, non-pneumonic LRTI (NP-LRTI) or no diagnosis of aLRTI. Pathogens were identified from standard-of-care and research microbiology. Outcomes were acute kidney injury (AKI), acute liver dysfunction, venous thromboembolism (VTE), in-hospital falls, reduced mobility at discharge, increased care requirements, and 30-day and 1-year mortality. Analyses were descriptive. Results Among 246,797 adult admissions, 21,456 aLRTD hospitalisations were included: 10,239 (47.7%) pneumonia, 7,742 (36.1%) NP-LRTI and 3,475 (16.2%) with no evidence of aLRTI. Of 19,152 tested aLRTD admissions, 8,503 (44.4%) had a positive microbiological/virological test, yielding 9,204 pathogen detections; 1,194 (6.2%) had co-infections, and SARS-CoV-2 was most frequent, with influenza the second most common in pneumonia and NP-LRTI. Pneumonia had greater severity than NP-LRTI and no diagnosis of aLRTI (median length of stay 6 vs 4 vs 4 days; ICU admission 3.4% vs 0.7% vs 0.5%, respectively). Overall, 22.2% developed AKI, 6.1% acute liver dysfunction, 0.6% DVT and 2.4% PE; 1.8% had a fall, 11.5% reduced mobility, and 16.6% required increased care at discharge. 30-day and 1-year mortality were highest for pneumonia (14.0% and 32.0%, respectively). Pathogen-specific analyses showed longer stays and higher complications and mortality rates for SARS-CoV-2 and Streptococcus pneumoniae, and shorter stays with lower complication and mortality rates for influenza and Haemophilus influenzae. Conclusions Non-cardiovascular complications and functional decline after aLRTI were common, particularly in pneumonic and SARS-CoV-2 or pneumococcal disease. These findings support routine surveillance for renal, hepatic, thromboembolic events, early mobilisation and rehabilitation, and consideration of multi-system outcomes when evaluating public health and economic value of vaccines and therapies.

10
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 1%
0.0%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

11
ECG-based longitudinal risk prediction across diseases and organ systems

ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.

2026-09-02 health informatics 10.64898/2026.08.29.26361697 medRxiv
Top 1%
0.0%
Show abstract

Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.

12
Certified large language model-based diagnostic decision support in rheumatology: the ALLIANCE multicentre randomised controlled trial

Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.

2026-09-02 rheumatology 10.64898/2026.08.29.26361715 medRxiv
Top 1%
0.0%
Show abstract

Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.

13
Optimizing Aqueous Humor Liquid Biopsy: Safety and Performance of a Short, Low-Dead-Space Ophthalmic Needle for Anterior Chamber Paracentesis

Singh, A. M.; Yeh, T.-C.; DeBoer, C.; Al-Moujahed, A.; Lin, J. B.; Smith, S. J.; Sanislo, S.; Janjua, K. A.; Lin, T.-C.; Almeida, D. R. P.; Mruthyunjaya, P.; Mahajan, V. B.

2026-09-02 ophthalmology 10.64898/2026.08.26.26361364 medRxiv
Top 1%
0.0%
Show abstract

Purpose: To evaluate the safety, procedural performance, sample recovery, and surgeon preference of an ophthalmic needle designed specifically for anterior chamber (AC) paracentesis. Methods: In this multicenter study, AC paracentesis was performed in clinic and operating-room settings using a 32-gauge x 4-mm needle with low dead space. The procedure was evaluated using a standardized physician survey. Prespecified outcomes included procedure-related adverse events (primary outcome), needle entry and handling, aspiration and sample recovery, comparative performance versus a 30-gauge needle, and physician preference for future use. Results: A total of 110 needle uses by eight surgeons were included. No ocular complications occurred, including lens or iris injury, hyphema, AC collapse, wound leak, hypotony, infection, or retinal complication, and no procedure required needle exchange or conversion to another device. Two technical events without ocular sequelae were noted, in which needle entry was partial thickness and did not reach the AC (1.8%; exact 95% CI, 0.2%-6.4%). Physicians rated needle entry, handling and sample recovery as good or excellent. Compared with a 30-gauge needle, the study needle was rated as at least comparable across all assessed domains. All surgeons rated it better or much better for intra-procedural safety and preferred it for future AC taps. Conclusions and Relevance: This short, 32-gauge low-dead-space ophthalmic needle demonstrated a favorable safety profile and was preferred over a 30-gauge needle by all surgeons. As aqueous humor liquid biopsy expands in clinical diagnostics and trials, an ophthalmic-specific needle design may help improve the consistency and safety of aqueous humor collection for molecular analysis and broader clinical use. Keywords: Anterior chamber paracentesis; Aqueous humor; Liquid biopsy; Low dead space; Ophthalmic needle

14
Evaluation of the Efficacy and Safety of Combination Therapy of Vamha and Myrha in the Management of PMOS: An Open-Label, Randomized, Multicentre, Comparative, Prospective Clinical Study

Patil, A.; Barathe, R.; Tate, D. M.; Kate, K.; Pande, S.; Gawande, N.; More, A.; Mahadik, S.; Berde, K.; Singhvi, R.

2026-09-02 sexual and reproductive health 10.64898/2026.08.20.26360875 medRxiv
Top 1%
0.0%
Show abstract

Introduction: Polyendocrine metabolic ovarian syndrome (PMOS), formerly known as polycystic ovary syndrome (PCOS), is a common endocrine disorder affecting women of reproductive age. Besides reproductive and metabolic disturbances, PMOS negatively impacts psychological well-being and quality of life. Despite available treatment options, there remains a need for safe and effective therapies that improve both clinical symptoms and fertility outcomes. Aim: To compare the efficacy of VAMHA and MYRHA tablet combination therapy with standard non-hormonal therapy in restoring regular menstruation. Secondary objectives included assessment of ovulation, menstrual symptoms, polycystic ovarian morphology, hormonal and metabolic parameters, anthropometric measures, and skin manifestations. Study Design: Open-label, randomized, multicentre, prospective comparative clinical study. Methods: Seventy-one women with PMOS were randomized to Group A (n=37) or Group B (n=34). Group A received VAMHA and MYRHA tablets (2 tablets each), while Group B received Metformin 500 mg plus Myoinositol 600 mg (1 tablet), twice daily for 180 days. Data were recorded in Case Report Forms. Statistical Analysis: Continuous variables were summarized using mean and standard deviation, while categorical variables were expressed as frequencies and percentages. Appropriate statistical tests, including Chi-square, were used. A p-value [&le;]0.05 was considered significant. Results: Significantly more participants in Group A achieved regular menstrual cycles than Group B (31 vs. 22; p<0.05). Ovulation occurred in 16 participants in Group A compared with 6 in Group B (p<0.05). Both groups showed significant improvement in menstrual irregularity and related symptoms. Significant reductions in Anti-Mullerian Hormone (AMH), fasting insulin, and body mass index (BMI) were observed in both groups (p<0.05). Resolution of polycystic ovarian morphology occurred in 13 participants (38.23%) in Group A and 10 (33.33%) in Group B. Both treatments were well tolerated with no major safety concerns. Conclusions: VAMHA and MYRHA combination therapy was superior to standard non-hormonal therapy in improving menstrual regularity and ovulation. It also produced favourable metabolic, hormonal, and ultrasonographic outcomes, suggesting its potential as a safe and effective option for comprehensive PMOS management and fertility enhancement.

15
Lung function trajectories in children with cystic fibrosis aged 3-17 years: impact of elexacaftor-tezacaftor-ivacaftor on lung function

Dyer, B. P.; Deery, M.; Heyman, R.; Robinson, P.; Wainwright, C.; Sly, P.; Ware, R.; Blake, T.

2026-09-02 respiratory medicine 10.64898/2026.08.31.26361791 medRxiv
Top 1%
0.0%
Show abstract

Background Elexacaftor-tezacaftor-ivacaftor (ETI) has been demonstrated to improve lung function in clinical trials; however, evidence describing effects on trajectories and whether long-term improvements are sustained (>1-year) is lacking. We estimated within-person lung clearance index (LCI) trajectories before and after ETI initiation, assessing changes in level and rate of change, alongside acute LCI change, up to three years after ETI initiation. Methods Prospective observational study of children at a tertiary hospital. Children aged 3-17 years with [&ge;]2 LCI testing occasions (i) before and (ii) after starting ETI were used to describe lung function trajectories. Children with [&ge;]1 pre-ETI and [&ge;]1 post-ETI LCI occasion(s) were used to describe acute LCI change after ETI initiation. Age-adjusted LCI trajectories for time periods (i) before and (ii) after ETI initiation were estimated using linear mixed-effects models, and pre- and post-ETI LCIs were compared using paired Wilcoxon tests. Results Mean pre-ETI and post-ETI longitudinal changes in LCI were -0.007 (95% CI: -0.28, 0.27; n=35) and 0.12 (95% CI: -0.17, 0.41; n=20) turnovers per year, respectively. Before ETI initiation, 57% (30/53) of patients had an LCI[&ge;]7.1 turnovers (indicating impaired lung function), compared to 26% (14/53) post-ETI, with a median LCI difference of -0.70 (95% CI -0.84, -0.46; p<0.001) turnovers. Within-individual variability in LCI decreased post-ETI. Conclusions Our real-world data within a unique longitudinal study provide a comprehensive picture of ETI benefit by outlining not only acute improvement in LCI but maintained stability in LCI trajectories and improved LCI stability sustained up to three years post-initiation.

16
Cross-System Meta-Analysis of Machine Learning Predictors Identifies Value-Specific Risk Drivers and Interactions Underlying Acute Kidney Injury

Chan, H. Y.; Li, D.; Yu, A. S. L.; Kellum, J. A.; Fuhrman, D. Y.; Xu, Q.; Chrischilles, E. A.; Cowell, L. G.; Chandaka, S.; Anzalone, A. J.; Kean, J.; McTigue, K. M.; Mosa, A. S. M.; Taylor, B.; Syed, M.; Waitman, L. R.; Hu, Y.; Liu, M.

2026-09-02 nephrology 10.64898/2026.08.31.26361849 medRxiv
Top 1%
0.0%
Show abstract

Background: Current understanding of acute kidney injury (AKI) risk factors remains largely descriptive, offering limited precision into how specific biomarker values or physiologic thresholds influence susceptibility. We aimed to synthesize knowledge from machine learning models trained across multiple health systems to identify generalizable, value-specific risk drivers and biomarker interactions contributing to AKI risk. Methods: We analyzed electronic health records (EHRs) from 785,497 adult inpatients between 2010 and 2019 across nine U.S. academic medical centers within PCORnet. Interpretable gradient boosting machine models were independently developed at each health system to quantify predictor-outcome associations. Meta-regression was applied to integrate these site-level results, characterize nonlinear value-risk relationships, and identify bivariate interactions between predictors. Results: Meta-analysis revealed consistent, value-specific risk drivers across health systems. An increase in glucose from 100 mg/dL to 140 mg/dL was associated with a 1.46-fold higher risk of AKI. Chloride and anion gap also demonstrated elevated AKI risk with risk increases overlapping portions of their reference ranges, with anion gap showing a 1.14-fold increase across 4-12 mmol/L and chloride a 1.28-fold increase across 96-100 mEq/L. Electrolytes including potassium, calcium, and sodium showed quadratic associations with AKI risk. Bivariate meta-regression identified interactions between key predictors, highlighting pathways that jointly modulate AKI risk. Conclusion: This cross-system meta-analysis synthesizes machine learning-derived evidence into clinically interpretable knowledge, revealing how specific biomarker ranges and interactions modulate AKI risk. By moving beyond surface-level associations to quantitative, generalizable physiologic thresholds, these findings provide actionable insights to enhance risk stratification and personalized prevention in hospital care.

17
GLP-1/GIP Uptake, Indication, and Access Pathways Among US Adults in the Understanding America Study

Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.

2026-09-02 endocrinology 10.64898/2026.08.28.26361368 medRxiv
Top 1%
0.0%
Show abstract

Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.

18
Social Determinants of Health in HIV/HBV Coinfection Compared with HIV and HBV Monoinfection: A Framework for Dynamic Social Vulnerability

Yendewa, G.; Chengsupanimit, T.; Dehghani, A.; Ahmed, A.; Mohareb, A.; Freeman, M.; Cohen, C.; Ofotokun, I.; Dube, K.

2026-09-02 hiv aids 10.64898/2026.08.31.26361856 medRxiv
Top 1%
0.0%
Show abstract

Human immunodeficiency virus (HIV) and hepatitis B virus (HBV) coinfection is associated with accelerated liver disease, but whether coinfection is associated with newly documented social determinants of health (SDoH) is unclear. We conducted a retrospective cohort study using TriNetX across 110 U.S. healthcare organizations (2010-2026). We propensity score matched adults with HIV/HBV to adults with HIV or HBV monoinfection. We organized newly documented SDoH indicators using a dynamic individual-level framework with four clinically recognized domains of social disadvantage: material vulnerability, healthcare access and engagement, interpersonal adversity, and psychosocial vulnerability. Matched cohorts included 10,071 HIV/HBV-HIV pairs and 9,659 HIV/HBV-HBV pairs (mean age, 47 years; 79% male; 66% non-White; median follow-up, 3.3 years). Over 178,900 person-years, HIV/HBV was associated with higher risk of the primary SDoH composite compared with HIV (11.5% vs 9.7%; incidence rate, 2.50 vs 1.97 per 100 person-years; hazard ratio [HR], 1.25; 95% confidence interval [CI], 1.15-1.37) and HBV (11.0% vs 6.4%; incidence rate, 2.39 vs 1.67; HR, 1.50; 95% CI, 1.35-1.67). HIV/HBV was also associated with higher material vulnerability and healthcare access and engagement composites in both comparisons, including housing instability, food insecurity, financial insecurity, insurance instability, and care disengagement/nonadherence (HR range, 1.22-3.33 vs HIV; 1.31-1.94 vs HBV). In the HBV comparison, HIV/HBV was additionally associated with interpersonal adversity, primary support stressors, and violence or victimization (HR range, 1.36-2.16). Findings were robust across sensitivity analyses. HIV/HBV was associated with more newly documented SDoH than monoinfection, supporting dynamic SDoH assessment.

19
Glaucoma and Diabetes Mellitus: A Comparative Evaluation of Comorbid Effect on Tear Quantity among Patients in Owerri, Imo State, Nigeria.

Chukwuoha, C. M.; Ovenseri-Ogbomo, G.; Azuamah, Y. C.; Odimegwu, N. E.; Obioma-Elemba, J. E.; Ugwoke, G.; Nkeremuzor, E. C.; Eronini, Y.; Ikoro, N. C.; Esenwah, E. C.

2026-09-02 ophthalmology 10.64898/2026.08.30.26361782 medRxiv
Top 1%
0.0%
Show abstract

Abstract Objective: Glaucoma is a chronic disorder that impairs ocular health and may exacerbate ocular surface disease leading to tear film instability, dry eye symptoms and decreased quality of life. This study compared changes in tear quantity among glaucoma subjects living with and without diabetes mellitus, attending an eye clinic in Nigeria. Methods: A comparative cross sectional research design was used. 157 subjects which comprised 74 glaucoma subjects living with diabetes mellitus and 83 glaucoma subjects living without diabetes mellitus participated in the study. Tear quantity assessment included the Schirmer I test and tear meniscus height (TMH) measurement. Descriptive statistics, independent samples t-test and Chi-square test were used to examine the data at 0.05 level of significance. Results: Glaucoma subjects living with diabetes mellitus showed substantially decreased tear production (11.4 +/- 6.8 mm) compared with glaucoma subjects living without diabetes mellitus (19.6 +/- 9.6 mm; p < 0.001). Tear meniscus height in glaucoma subjects living with diabetes mellitus (0.8 +/- 0.3 mm) was significantly greater than in subjects living without diabetes mellitus (0.7 +/- 0.3 mm; p = 0.034). Conclusion: Diabetes mellitus dramatically deteriorates the ocular surface function in glaucoma subjects by decreasing tear production, altering the tear meniscus height and increasing the severity of ocular surface symptoms. Routine glaucoma care, especially in patients with diabetes mellitus, should include a full ocular surface evaluation including Schirmer I test, TBUT, TMH, and OSDI assessment to allow early detection and management of ocular surface disease, better treatment adherence, and improved visual outcomes. Keywords: Glaucoma, Diabetes Mellitus, Tear production, Tear Meniscus Height, Ocular Surface Disease.

20
Non-inferior survival and enhanced longevity with initial low-dose versus full-dose enzalutamide: a single-centre real-world prostate cancer study

Gorobets, O.; Vinh-Hung, V.

2026-09-02 oncology 10.64898/2026.08.28.26361616 medRxiv
Top 1%
0.0%
Show abstract

Background: Prostate cancer enzalutamide treatment is approved at a standard dose of 160 mg daily. Concerns for real-world patients -- older and more fragile than those enrolled in clinical trials -- have prompted consideration of initiating treatment with lower doses, but the long-term efficacy of this approach remains unknown. We evaluate the long-term survival and longevity in patients treated with standard versus upfront low-dose enzalutamide. Methods: Retrospective analysis of 151 patients treated with enzalutamide (102 receiving 160 mg; 49 receiving [&le;]80 mg) between 2014--2021 at the Centre Hospitalier Universitaire de Martinique, with complete follow-up through end of life (98.7% completeness of follow-up). Primary outcomes were overall survival (OS), progression-free survival (PFS), and longevity (attained age). Results: Doses [&le;]80 mg were associated with longer median OS (36.3 vs. 20.7 months), improved restricted mean OS (difference of 0.7 years, p=0.05), and enhanced longevity (median 82.5 vs. 78.3 years, p=0.004). PSA response rate at 12 weeks was higher with lower-dose (71.4% vs. 48.8%, p=0.016). In multivariable models adjusted for prognostic factors, [&le;]40 mg compared with 160 mg was non-inferior regarding OS (HR=0.61, 95% CI 0.36--1.06), superior regarding PFS (HR=0.59, 95% CI 0.35--0.99), and superior regarding longevity (HR=0.48, 95% CI 0.28--0.84). Bone metastasis, poor performance status, PSA response, time to PSA nadir, and disease duration were independent predictors of outcomes. A post-hoc analysis revealed a strong association between dose and physician-prescribing profiles, ranging from "endorse-lowest-dose" to "never-deviate-from-full-dose". Conclusions: Lower doses of enzalutamide were non-inferior to full-dose. Dose-adapted strategies warrant further investigation.